Papers with data-to-text generation
On Training Instance Selection for Few-Shot Neural Text Generation (2021.acl-short)
Copied to clipboard
| Challenge: | Pretraining large neural networks with a language modeling objective has led to dramatic improvements in text generation. |
| Approach: | They propose a selection strategy to select few-shot training instances based on unlabeled data to identify the most worthwhile data points that should be annotated under some budget of labeling cost. |
| Outcome: | The proposed strategy outperforms random sampling on three text generation tasks. |
Deep Learning Approaches to Text Production (N18-6)
Copied to clipboard
| Challenge: | Text production is a key component of many NLP applications . Claire Gardent is based in France and is pursuing research in text production . |
| Approach: | This tutorial will cover the fundamentals and state-of-the-art research on neural models for text production. |
| Outcome: | This tutorial will cover the fundamentals and the state-of-the-art research on neural models for text production. |
WikiTableT: A Large-Scale Data-to-Text Dataset for Generating Wikipedia Article Sections (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets for data-to-text generation focus on single-sentence generation or long-form generation. |
| Approach: | They create a dataset that pairs Wikipedia sections with tabular data and various metadata. |
| Outcome: | The proposed dataset can generate fluent and high quality texts but struggle with coherence and factuality. |
Data-to-text Generation with Macro Planning (2021.tacl-1)
Copied to clipboard
| Challenge: | Recent approaches to data-to-text generation adopt the encoder-decoder architecture . however, these models perform poorly at selecting appropriate content and ordering it coherently . |
| Approach: | They propose a neural model with a macro planning stage followed by a generation stage . they use data from databases of records, simulations of physical systems, accounting spreadsheets . |
| Outcome: | The proposed model outperforms baselines on two data-to-text benchmarks . it uses the encoderdecoder architecture and is compared with existing models . |
SPOR: A Comprehensive and Practical Evaluation Method for Compositional Generalization in Data-to-Text Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies on compositional generalization in data-to-text generation focus on one manifestation, Systematicity, Productivity, Order invariance, and Rule learnability. |
| Approach: | They propose a method for evaluation of compositional generalization in data-to-text generation that includes four aspects of manifestations and allows high-quality evaluation without additional manual annotations. |
| Outcome: | The proposed method is based on two datasets and evaluates existing language models including LLMs. |
Data-to-text Generation with Variational Sequential Planning (2022.tacl-1)
Copied to clipboard
| Challenge: | Recent advances in data-to-text generation have greatly facilitated the task of generating textual output from non-linguistic input. |
| Approach: | They propose a neural model enhanced with a planning component responsible for organizing high-level information in a coherent and meaningful way. |
| Outcome: | The proposed model outperforms baseline models and is sample-efficient in the face of limited training data. |
TabGenie: A Toolkit for Table-to-Text Generation (2023.acl-demo)
Copied to clipboard
| Challenge: | TabGenie enables researchers to explore, preprocess, and analyze data-to-text generation datasets. |
| Approach: | They present TabGenie, a toolkit which enables researchers to explore, preprocess, and analyze a variety of data-to-text generation datasets. |
| Outcome: | The toolkit provides an interactive mode for debugging table-to-text generation, side-by-side comparison of generated system outputs, and easy exports for manual analysis. |
Structured Discourse Representation for Factual Consistency Verification (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to verify factual consistency of text capture a performance gap compared with sentence-level entailment. |
| Approach: | They propose a method that combines structured discourse information extraction with a classifier, FDSpotter, for factual consistency verification. |
| Outcome: | The proposed method achieves competitive performance on two tasks: data-to-text generation and text summarisation. |
Neural data-to-text generation: A comparison between pipeline and end-to-end architectures (D19-1)
Copied to clipboard
| Challenge: | Traditionally, data-to-text applications have been designed using a modular pipeline architecture, in which the non-linguistic input data is converted into natural language through several intermediate transformations. |
| Approach: | They propose to use Gated-Recurrent Units and Transformer to implement neural pipelines for data-to-text generation. |
| Outcome: | The proposed models generalize better to unseen inputs and have better performance than the existing pipeline architectures. |
MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance (D19-1)
Copied to clipboard
| Challenge: | Existing evaluation metrics are not capable of evaluating text quality. |
| Approach: | They propose a metric that compares system output against reference texts based on semantics rather than surface forms. |
| Outcome: | The proposed metric shows a high correlation with human judgment of text quality on a number of text generation tasks. |
Does the Order of Training Samples Matter? Improving Neural Data-to-Text Generation with Curriculum Learning (2021.eacl-main)
Copied to clipboard
| Challenge: | Recent advances in data-to-text generation have been focused on curriculum learning, which is a process of presenting training data in a specific order, starting from easy examples and moving on to more difficult ones, as the learner becomes more competent. |
| Approach: | They propose to use a curriculum learning process to change the order of training samples in a model based on the model's competence to improve model performance and convergence speed. |
| Outcome: | The proposed model shows faster convergence speed and reduced training time by 38.7% and performance by 4.84 BLEU. |
Neural Data-to-Text Generation with LM-based Text Augmentation (2021.eacl-main)
Copied to clipboard
| Challenge: | Neural data-to-text generation is a difficult task for many new applications because of a lack of training data. |
| Approach: | They propose a few-shot approach that augments the data available for training by generating new text samples based on replacing specific values by alternative ones from the same category and pairing the new text with data samples. |
| Outcome: | The proposed approach outperforms fully supervised sequence-to-sequence models with less than 10% of the training set on both datasets. |
Plan-then-Generate: Controlled Data-to-Text Generation via Planning (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies focus on producing results that are close to the references, i.e. what to generate and in what order (the output structure) cannot be explicitly controlled by the users. |
| Approach: | They propose a Plan-then-Generate framework to improve the controllability of neural data-to-text models. |
| Outcome: | The proposed model can control both the intra-sentence and inter-sentent structure of the generated output. |
Point Precisely: Towards Ensuring the Precision of Data in Generated Texts Using Delayed Copy Mechanism (C18-1)
Copied to clipboard
| Challenge: | Recent neural generation systems have shown significant progress on data-to-text generation tasks. |
| Approach: | They propose a two-stage approach with a delayed copy mechanism to improve the precision of data records in the generated texts. |
| Outcome: | The proposed approach improves the accuracy of the generated texts on a RotoWire dataset. |
High-quality Data-to-Text Generation for Severely Under-Resourced Languages with Out-of-the-box Large Language Models (2024.findings-eacl)
Copied to clipboard
| Challenge: | Pretrained large language models (LLMs) can bridge the performance gap for under-resourced languages by substantial margins, as measured by both automatic and human evaluations. |
| Approach: | They propose to use pretrained large language models to bridge this gap by automating and evaluating data-to-text generation in under-resourced languages. |
| Outcome: | The proposed model can set the state of the art for under-resourced languages by substantial margins, as measured by both automatic and human evaluations. |
Open Domain Question Answering with A Unified Knowledge Interface (2022.acl-long)
Copied to clipboard
| Challenge: | a retriever-reader framework is popular for open domain question answering . however, accessing heterogeneous knowledge sources through a unified interface remains unknown . |
| Approach: | They propose a retriever-reader framework that uses explicit knowledge to access heterogeneous knowledge sources through a unified interface. |
| Outcome: | The proposed framework can benefit from the expanded knowledge index, the authors show . their approach sets the single-model state-of-the-art on Natural Questions . |
Generating Syntactic Paraphrases (D18-1)
Copied to clipboard
| Challenge: | Using data-to-text generation, text-totext generation and text reduction, we show that conditioning text generation on syntactic constraints permits the generation of syntakically distinct paraphrases for the same input. |
| Approach: | They propose to use four different models for automatic generation of syntactic paraphrases to study the automatic generation process. |
| Outcome: | The proposed models can generate syntactic paraphrases for the same input and exploit different types of input to increase the number of distinct paraphrased for a given input. |
Data-to-Text Generation with Style Imitation (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent approaches to data-to-text generation focus on improving content fidelity, but lack explicit control over writing styles. |
| Approach: | They propose a way to control writing styles by using existing sentences as "soft" templates . they conduct experiments in restaurants and sports domains to test their approach . |
| Outcome: | The proposed approach achieves stronger performance than a range of comparison methods. |
Grouped-Attention for Content-Selection and Content-Plan Generation (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Recent neural data-to-text generation models explicitly learn content-plan given a set of attributes as input. |
| Approach: | They propose a neural content-planner that captures local and global contexts . they use a token-level attention constrained within each input attribute . |
| Outcome: | The proposed model outperforms competitors by 4.92%, 4.70%, and 16.56% on real-world datasets. |
Data-to-text Generation with Entity Modeling (P19-1)
Copied to clipboard
| Challenge: | Recent approaches to data-to-text generation have shown great promise thanks to the use of large-scale datasets and the application of neural network architectures which are trained end-to end. |
| Approach: | They propose an entity-centric neural architecture for data-to-text generation which uses hierarchical attention to create entity-specific representations which are dynamically updated. |
| Outcome: | The proposed model outperforms baselines in automatic and human evaluation on the RotoWire benchmark and a five-times larger dataset on the baseball domain. |
On Hallucination and Predictive Uncertainty in Conditional Language Generation (2021.eacl-main)
Copied to clipboard
| Challenge: | Modern deep neural network models have brought drastic improvements in generation quality measured by standard metrics on different natural language generation tasks. |
| Approach: | They propose a beam search extension to reduce hallucination in conditional language generation by adding a prediction extension to beam search. |
| Outcome: | The proposed extension improves trading performance on standard metric for less hallucination with the proposed beam search variant. |
Text Generation with Exemplar-based Adaptive Decoding (N19-1)
Copied to clipboard
| Challenge: | Empirical results show that the proposed model achieves strong performance and outperforms comparable baselines. |
| Approach: | They propose a conditioned text generation model that uses a template-based approach to generate content from input text. |
| Outcome: | The proposed model outperforms baselines on abstractive text summarization and data-to-text generation. |
Learning Semantic Correspondences from Noisy Data-text Pairs by Local-to-Global Alignments (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for data-to-text generation use a large-scale training corpus to learn semantic correspondences between structured input data and associated texts. |
| Approach: | They propose a local-to-global alignment framework that uses local and global models to learn semantic correspondences from large-scale datasets. |
| Outcome: | The proposed framework can be generalized to restaurant and computer domains and improve alignment accuracy. |
Enhancing Neural Data-To-Text Generation Models with External Background Knowledge (D19-1)
Copied to clipboard
| Challenge: | Recent neural models for data-to-text generation rely on parallel pairs of data and text to learn writing knowledge. |
| Approach: | They propose to enhance neural models with external knowledge to improve fidelity of generated text. |
| Outcome: | The proposed model improves on Wikipedia infobox-to-text datasets on 21 datasets. |
Long and Diverse Text Generation with Planning-based Hierarchical Variational Model (D19-1)
Copied to clipboard
| Challenge: | Existing methods for data-to-text generation are insufficient to produce long and diverse texts. |
| Approach: | They propose a planning-based hierarchical variational model that plans a sequence of groups and then realizes each sentence conditioned on the planning result and the previously generated context. |
| Outcome: | The proposed model outperforms state-of-the-art models in long and diverse text generation. |
PixT3: Pixel-based Table-To-Text Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Table-to-text generation is a visual recognition task that uses textual descriptions from structured inputs. |
| Approach: | They propose to rethink data-to-text generation as a visual recognition task by removing the need for rendering the input in a string format. |
| Outcome: | The proposed model overcomes the challenges of linearization and input size limitations and is applicable to open-ended and controlled generation settings. |
Generating Textual Explanations for Machine Learning Models Performance: A Table-to-Text Task (2022.lrec-1)
Copied to clipboard
| Challenge: | Numerical tables are widely used to communicate or report the classification performance of machine learning models with respect to a set of evaluation metrics. |
| Approach: | They propose a task where neural models are trained to generate textual explanations based on the metrics’ scores reported in numerical tables. |
| Outcome: | The proposed model outperforms existing methods and can be used to explain the performance of ML models. |
Prompt Optimization via Adversarial In-Context Learning (2024.acl-long)
Copied to clipboard
Do Long, Yiran Zhao, Hannah Brown, Yuxi Xie, James Zhao, Nancy Chen, Kenji Kawaguchi, Michael Shieh, Junxian He
| Challenge: | Existing methods to optimize prompts for in-context learning are based on adversarial learning and are computationally efficient and extensible to other LLMs and tasks. |
| Approach: | They propose a method to optimize prompts for in-context learning by a generator and a discriminator. |
| Outcome: | The proposed method improves state-of-the-art prompt optimization techniques on 13 generation and classification tasks including summarization, arithmetic reasoning, machine translation, data-to-text generation, and the MMLU and big-bench hard benchmarks. |
Operation-guided Neural Networks for High Fidelity Data-To-Text Generation (D18-1)
Copied to clipboard
| Challenge: | Recent neural models for data-to-text generation generate descriptions that are not consistent with structured data. |
| Approach: | They propose a framework for data-to-text generation that uses symbolic operations to generate texts from structured data. |
| Outcome: | The proposed framework improves the fidelity of the generated texts to the input structured data. |
Improving Encoder by Auxiliary Supervision Tasks for Table-to-Text Generation (2021.acl-long)
Copied to clipboard
| Challenge: | Experimental results show that our method not only has a good generalization but also outperforms previous methods on several metrics: BLEU, Content Selection, Content Ordering. |
| Approach: | They propose to build an entity graph from the input tables and introduce a reasoning module to perform reasoning on the graph. |
| Outcome: | The proposed method outperforms previous methods on several metrics: BLEU, Content Selection, Content Ordering. |
Table-To-Text generation and pre-training with TabT5 (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) are limited when it comes to structured or semi-structured domains like tables. |
| Approach: | They propose an encoder-decoder model that generates natural language text based on tables and textual inputs. |
| Outcome: | TabT5 achieves 15% increase in sequence accuracy on spreadsheet formula prediction and data-to-text generation domains. |
Selective Token Generation for Few-shot Natural Language Generation (2022.coling-1)
Copied to clipboard
| Challenge: | Experimental results show that the proposed selective token generation algorithm outperforms the previous additive learning algorithms based on the PLMs. |
| Approach: | They propose an additive learning algorithm that selectively outputs language tokens between a task-general PLM and a specific adapter during training and inference. |
| Outcome: | The proposed algorithm outperforms existing methods on few-shot natural language generation tasks. |
Time-aware Prompting for Text Generation (2022.findings-emnlp)
Copied to clipboard
| Challenge: | a new study investigates the effects of incorporating timestamps into generation systems . textual prompts focus more on non-temporal information and are less sensitive to given timestams . |
| Approach: | They propose a data-to-text generation dataset that includes chronologically ordered revisions of biographical articles from English Wikipedia. |
| Outcome: | The proposed models improve the quality of the data-to-text generation dataset TempWikiBio . the proposed models are more sensitive to time-aware prompts than textual prompts . |
Grounded Keys-to-Text Generation: Towards Factual Open-Ended Generation (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Large pre-trained language models have enabled open-ended generation frameworks to tackle a variety of tasks beyond data-to-text generation. |
| Approach: | They propose a new task to generate a factual description about an entity given guiding keys and grounding passages using a dataset. |
| Outcome: | The proposed model improves factual correctness and recall significantly compared to previous models. |
Pruning Pre-trained Language Models with Principled Importance and Self-regularization (2023.findings-acl)
Copied to clipboard
| Challenge: | Pre-trained language models often contain a vast amount of parameters, posing nontrivial requirements for storage and computation. |
| Approach: | They propose a pruning method where model prediction is regularized by the latest checkpoint with increasing sparsity throughout pruning. |
| Outcome: | The proposed approach is effective at sparsity levels, and can be applied to natural language understanding, question answering, and data-to-text generation tasks. |
DORB: Dynamically Optimizing Multiple Rewards with Bandits (2020.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in end-to-end neural networks-based approaches have shown wide success in sequence generation tasks. |
| Approach: | They propose to optimize multiple metric rewards simultaneously using a multi-armed bandit approach . they empirically show the effectiveness of their approaches via various automatic metrics and human evaluation . |
| Outcome: | The proposed approach improves on question generation and data-to-text generation using a bandit approach. |
TLM: Token-Level Masking for Transformers (2023.emnlp-main)
Copied to clipboard
| Challenge: | Structured dropout approaches have been investigated to regularize the multi-head attention mechanism in Transformers. |
| Approach: | They propose a new regularization scheme based on token-level rather than structure-level to reduce overfitting by manipulating the connections between tokens in the multi-head attention via masking. |
| Outcome: | The proposed regularization scheme outperforms attention dropout and DropHead on 18 datasets and can establish a new record on the data-to-text benchmark Rotowire (18.93 BLEU). |
Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning (2023.acl-long)
Copied to clipboard
Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma, Patrick Ng, Zhiguo Wang, Bonan Min, William Yang Wang, Kathleen McKeown, Vittorio Castelli, Dan Roth, Bing Xiang
| Challenge: | Existing methods for data-to-text generation focus on specific types of structured data. |
| Approach: | They propose a method that provides a unified representation that can handle various forms of structured data such as tables, knowledge graph triples, and meaning representations. |
| Outcome: | The proposed method improves zero-shot and few-shot scenarios and can adapt to new structured data. |
LLM Multi-Agent Systems for Long Triple Set Data-to-Text Generation (2026.findings-acl)
Copied to clipboard
Chinonso Cynthia Osuji, Simon Mille, Mark Andrade, Jane Adkins, Ornait O’Connell, Elaine Uí Dhonnchadha, Bláithín Heffernan, Fírinne Nic an tSaoir, Anya Belz, Thiago Castro Ferreira, Brian Davis
| Challenge: | Existing data-to-text benchmarks that do not involve content selection feature short input-output pairs designed for sentence or paragraph-level generation with reference texts spanning only a few dozen tokens. |
| Approach: | They propose a system that generates multi-paragraph outputs in English and Irish . they compare a multi-agent configuration against a single-task variant . |
| Outcome: | The proposed framework generates multi-paragraph outputs in English and Irish . human evaluation and LLM-as-a-judge score better in both languages . |